Papers by Kiet Van Nguyen
ViSoLex: An Open-Source Repository for Vietnamese Social Media Lexical Normalization (2025.coling-demos)
Copied to clipboard
| Challenge: | ViSoLex is an open-source repository for Vietnamese lexical normalization . it provides two core services: Non-Standard Word (NSW) Lookup and Lexical Normalization enabling users to retrieve standard forms of informal language and standardize text containing NSWs. |
| Approach: | They propose to integrate pre-trained language models and weakly supervised learning techniques to ensure accurate and efficient normalization. |
| Outcome: | The system provides two core services: Non-Standard Word (NSW) Lookup and Lexical Normalization, enabling users to retrieve standard forms of informal language and standardize text containing NSWs. |
ViHOS: Hate Speech Spans Detection for Vietnamese (2023.eacl-main)
Copied to clipboard
| Challenge: | Increasing use of social networking sites can cause problems for human moderators to review tagged comments. |
| Approach: | They present a dataset that contains 26k spans on 11k comments and detailed annotation guidelines . they also provide definitions of hateful and offensive spans in Vietnamese comments . |
| Outcome: | The proposed dataset shows that it is difficult to detect specific types of spans in the dataset . the dataset is the first human-annotated corpus containing 26k spans on 11k comments . |
Revealing Weaknesses of Vietnamese Language Models Through Unanswerable Questions in Machine Reading Comprehension (2023.eacl-srw)
Copied to clipboard
| Challenge: | Existing problems in Vietnamese Machine Reading Comprehension systems are limited due to multilinguality, which limits the ability of multilingual models to develop state-of-the-art systems. |
| Approach: | They propose to modify the process of annotating unanswerable questions to improve the quality of unanswered questions to a higher level of difficulty for Machine Reading Comprehension systems to solve. |
| Outcome: | The proposed modification improves the quality of unanswerable questions to a higher level of difficulty for Machine Reading Comprehension systems to solve. |
ViGoEmotions: A Benchmark Dataset For Fine-grained Emotion Detection on Vietnamese Texts (2026.eacl-long)
Copied to clipboard
| Challenge: | Recent advances in NLP have greatly improved outcomes in emotion prediction and harmful content detection. |
| Approach: | They propose to classify Vietnamese comments into 27 distinct emotions using a model-based lexical normalization system and a transformer-based model. |
| Outcome: | The proposed corpus of 20,664 social media comments is based on a novel model that can support multiple architectures, but its quality and preprocessing strategies remain key factors influencing performance. |
A Large-Scale Benchmark for Vietnamese Sentence Paraphrases (2025.findings-naacl)
Copied to clipboard
| Challenge: | 1.2M original–paraphrase pairs were generated using a hybrid approach to generate high-quality paraphrases. |
| Approach: | They present a high-quality Vietnamese dataset for sentence paraphrasing . they used automatic paraphrase generation and manual evaluation to ensure high quality . |
| Outcome: | The proposed dataset is the first large-scale study on Vietnamese paraphrasing . it combines automatic paraphrase generation with manual evaluation to ensure high quality . |
ViNLI: A Vietnamese Corpus for Studies on Open-Domain Natural Language Inference (2022.coling-1)
Copied to clipboard
| Challenge: | a large-scale corpus is needed for studies on natural language inference (NLI) for Vietnamese, which can be considered a low-resource language. |
| Approach: | They propose a corpus for evaluating Vietnamese natural language inference models . they use a human-annotated corpus extracted from more than 800 online news articles . |
| Outcome: | The ViNLI corpus is created and evaluated with a strict process of quality control . the best system performance is still far from human performance (a 14.20% gap in accuracy). |